Skip to content

[AMD][AgentX] DSv4 MI355X agentic disagg: add EAGLE/MTP speculative-decoding arm / DSv4 MI355X 智能体分离式推理:新增 EAGLE/MTP 投机解码分支 - #2309

Open
ichbinblau wants to merge 15 commits into
mainfrom
amd/agentx-v1.0-th-hicon-mtp
Open

[AMD][AgentX] DSv4 MI355X agentic disagg: add EAGLE/MTP speculative-decoding arm / DSv4 MI355X 智能体分离式推理:新增 EAGLE/MTP 投机解码分支#2309
ichbinblau wants to merge 15 commits into
mainfrom
amd/agentx-v1.0-th-hicon-mtp

Conversation

@ichbinblau

@ichbinblau ichbinblau commented Jul 23, 2026

Copy link
Copy Markdown
Collaborator

Summary

DSv4 FP4 MI355X SGLang PD-disaggregated agentic-coding benchmark: HiCache sizing/NCCL/router fixes, the TP8/EP8/DPA hicache search-space arm, and a new EAGLE/MTP speculative-decoding arm. Originally split as #2308 (fixes) stacked under #2309 (MTP), now combined into a single PR targeting main directly.

What's changed

HiCache sizing (ratio vs size)

  • dsv4_fp4_mi355x_sglang-disagg.sh: set HICACHE_RATIO=4 and FORCE_HICACHE_RATIO=1 so this recipe explicitly opts into ratio-based sizing (DeepSeek-V4's hybrid HiCache pool rejects --hicache-size).
  • job.slurm: forward FORCE_HICACHE_RATIO into the container's HiCache env file (hicache_mc_<job>.env) alongside the other HICACHE_* vars -- it was previously silently dropped, so the override never reached the server.

NCCL / decode cuda-graph crash workaround

RDMA QoS/DCQCN + GPU-sanity pre-flight gates (job.slurm)

  • New helpers/rdma_check.sh: validates NIC QoS (PFC/DSCP) and DCQCN configuration on every allocated node via srun before any container/GPU time is spent. Override with SKIP_RDMA_CHECK=1.
  • New helpers/gpu_sanity.sh: after container pre-clean/teardown, fails fast (naming the offending process) if a bare, non-containerized process is still holding GPU VRAM.

Router policy fix + DPA (ep=8/dp-attn) hicache arm

  • server_sglang.sh (agentic branch only -- DI branch untouched): fixes a bug where the script read ROUTER_PREFILL_POLICY but recipe scripts export PREFILL_ROUTER_POLICY, so the recipe's policy override was silently ignored and the router always fell back to the hardcoded default. Also switches the default prefill router policy cache_aware -> consistent_hashing, adds --dp-aware to the router flags, and comments out the separate --decode-policy override (superseded by --dp-aware).
  • Adds PREFILL_DP_ATTN/DECODE_DP_ATTN to the env-var whitelist forwarded into the sibling client container, alongside the existing PREFILL_ENABLE_DP/DECODE_ENABLE_DP.
  • configs/amd-master.yaml: bump dsv4-fp4-mi355x-sglang-disagg-agentic-hicache image to v0.5.15.post1-rocm720-mi35x-20260719; re-enable (previously commented out) the TP8/EP8/DPA (ep=8, dp-attn=true) hicache search-space arm at conc-list: [64, 96, 128].
  • trace_replay.sh: when PREFILL_ENABLE_DP=true, export AIPERF_HTTP_X_SMG_ROUTING_KEY_FROM_CORRELATION_ID=true so aiperf's HTTP routing key is derived from the correlation ID under DP-attention.

EAGLE/MTP speculative-decoding arm

  • models.yaml: add mtp_flags: "--speculative-algorithm EAGLE --speculative-eagle-topk 1" to DeepSeek-V4-Pro-AgentX. DeepSeek-V4's MTP head is a native NextN/EAGLE draft shipped with the model, so SGLang defaults the draft path to the main checkpoint -- no --speculative-draft-model-path needed. --speculative-num-steps/--speculative-num-draft-tokens are already derived from DECODE_MTP_SIZE in server_sglang.sh; mtp_flags is only appended to the decode server config (MTP is decode-only).
  • configs/amd-master.yaml: new dsv4-fp4-mi355x-sglang-disagg-agentic-hicache-mtp config-key with two search-space arms mirroring the non-MTP hicache config's topologies (spec-decoding: "mtp", DECODE_MTP_SIZE=3): TP8/EP1 no-DP at conc-list: [2, 4, 8, 16, 32], and TP8/EP8/DPA at conc-list: [64, 96, 128].

perf-changelog.yaml

Two appended entries (tail of file, per the append-only convention), one per config-key touched, both linked to this PR: dsv4-fp4-mi355x-sglang-disagg-agentic-hicache (image bump / HiCache / NCCL / router fixes / DPA arm) and dsv4-fp4-mi355x-sglang-disagg-agentic-hicache-mtp (new EAGLE/MTP arm).

Test plan

  • bash -n on all touched shell scripts (env.sh, job.slurm, server_sglang.sh, trace_replay.sh, setup_deps.sh, helpers/gpu_sanity.sh, helpers/rdma_check.sh) -- syntax OK
  • configs/amd-master.yaml, benchmarks/multi_node/amd_utils/models.yaml, perf-changelog.yaml -- YAML parses cleanly
  • generate_sweep_configs.py full-sweep resolves all 16 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache* configs correctly: 8 non-MTP (TP8/EP1 conc 2/4/8/16/32, TP8/EP8/DPA conc 64/96/128) + 8 MTP twins with identical topologies plus spec-decoding: mtp
  • Diffed against main to confirm no unintended regression in the IS_AGENTIC/DI router split or the router-readiness canary (wait_or_die)
  • CI: e2e Test - dsv4-fp4-mi355x-sglang-disagg-agentic-hicache and ...-hicache-mtp (multi-node, mi355x-amds)

中文说明

DSv4 FP4 MI355X SGLang 预填充/解码分离式(disaggregated)智能体编码(agentic-coding)基准测试:修复 HiCache 容量配置、NCCL 以及路由策略问题,重新启用 TP8/EP8/DPA 的 hicache 搜索空间分支,并新增一个 EAGLE/MTP 投机解码(speculative decoding)分支。原本拆分为 #2308(修复项)叠加在 #2309(MTP)之上,现已合并为一个直接面向 main 的单一 PR。

HiCache 容量配置(ratio 与 size)

  • dsv4_fp4_mi355x_sglang-disagg.sh:设置 HICACHE_RATIO=4FORCE_HICACHE_RATIO=1,使该配方显式启用基于比例(ratio-based)的容量配置(DeepSeek-V4 的混合 HiCache 池不支持 --hicache-size)。
  • job.slurm:将 FORCE_HICACHE_RATIO 转发到容器的 HiCache 环境变量文件(hicache_mc_<job>.env)中,此前该变量被静默丢弃,导致覆盖配置无法生效。

NCCL / 解码 CUDA-graph 崩溃的规避方案

  • env.sh:新增 TORCH_NCCL_BLOCKING_WAIT/NCCL_BLOCKING_WAIT=1,以规避 ROCm 7.2 在解码 CUDA-graph 捕获阶段出现的 HIP 运行时看门狗(watchdog)竞态问题;NCCL_DEBUG=INFO 保持注释状态以减少日志噪音。
  • models.yaml:为 DeepSeek-V4-Pro-AgentX 禁用预填充 CUDA-graph(disable_cuda_graph: true),从 dp_flags 中移除 --enable-two-batch-overlap,将预填充 mem_fraction_static 从 0.8 降至 0.72,非 DP 场景下 max_running_requests 从 128 降至 64;移除解码侧多余的 disable_cuda_graph: false(已是代码默认值)。
  • 镜像版本升级:lmsysorg/sglang-rocm:v0.5.14-rocm720-mi35x-20260710 -> v0.5.15.post1-rocm720-mi35x-20260719

RDMA QoS/DCQCN 与 GPU 健康检查前置校验(job.slurm

  • 新增 helpers/rdma_check.sh:在占用任何容器/GPU 时间之前,通过 srun 校验每个分配节点的网卡 QoS(PFC/DSCP)与 DCQCN 配置;可通过 SKIP_RDMA_CHECK=1 跳过。
  • 新增 helpers/gpu_sanity.sh:容器预清理/收尾之后,若仍有裸机(非容器化)进程占用 GPU 显存,则快速失败并报告具体进程。

路由策略修复 + DPA(ep=8/dp-attn)hicache 分支

  • server_sglang.sh(仅智能体分支,DI 分支未改动):修复脚本读取 ROUTER_PREFILL_POLICY 而配方实际导出的是 PREFILL_ROUTER_POLICY 的问题,此前该覆盖配置被静默忽略,路由器始终回退到硬编码默认值。同时将默认预填充路由策略由 cache_aware 切换为 consistent_hashing,为路由参数新增 --dp-aware,并注释掉原有的 --decode-policy 覆盖项(已被 --dp-aware 取代)。
  • PREFILL_DP_ATTN/DECODE_DP_ATTN 加入转发到客户端容器的环境变量白名单,与现有的 PREFILL_ENABLE_DP/DECODE_ENABLE_DP 并列。
  • configs/amd-master.yaml:将 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache 镜像升级至 v0.5.15.post1-rocm720-mi35x-20260719;重新启用此前被注释的 TP8/EP8/DPA(ep=8, dp-attn=true)hicache 搜索空间分支,并发列表为 [64, 96, 128]
  • trace_replay.sh:当 PREFILL_ENABLE_DP=true 时,导出 AIPERF_HTTP_X_SMG_ROUTING_KEY_FROM_CORRELATION_ID=true,使 aiperf 的 HTTP 路由键在 DP-attention 下基于关联 ID(correlation ID)派生。

EAGLE/MTP 投机解码分支

  • models.yaml:为 DeepSeek-V4-Pro-AgentX 新增 mtp_flags: "--speculative-algorithm EAGLE --speculative-eagle-topk 1"。DeepSeek-V4 的 MTP 头是模型自带的原生 NextN/EAGLE 草稿模型,因此 SGLang 默认将草稿路径指向主检查点,无需 --speculative-draft-model-path--speculative-num-steps/--speculative-num-draft-tokens 已在 server_sglang.sh 中由 DECODE_MTP_SIZE 派生;mtp_flags 仅追加到解码侧服务器配置中(MTP 仅作用于解码)。
  • configs/amd-master.yaml:新增 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache-mtp 配置项,包含两个搜索空间分支,拓扑结构与非 MTP 版本的 hicache 配置一致(spec-decoding: "mtp"DECODE_MTP_SIZE=3):TP8/EP1 非 DP,并发列表为 [2, 4, 8, 16, 32];以及 TP8/EP8/DPA,并发列表为 [64, 96, 128]

perf-changelog.yaml

在文件末尾追加两条记录(遵循只追加约定),分别对应本次改动涉及的两个配置项,均关联本 PR:dsv4-fp4-mi355x-sglang-disagg-agentic-hicache(镜像升级/HiCache/NCCL/路由修复/DPA 分支)与 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache-mtp(新增 EAGLE/MTP 分支)。

测试计划

  • 对所有改动的 shell 脚本(env.shjob.slurmserver_sglang.shtrace_replay.shsetup_deps.shhelpers/gpu_sanity.shhelpers/rdma_check.sh)执行 bash -n —— 语法正确
  • configs/amd-master.yamlbenchmarks/multi_node/amd_utils/models.yamlperf-changelog.yaml —— YAML 解析无误
  • generate_sweep_configs.py full-sweep 正确解析全部 16 个 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache* 配置:8 个非 MTP(TP8/EP1 并发 2/4/8/16/32,TP8/EP8/DPA 并发 64/96/128)+ 8 个拓扑相同并带 spec-decoding: mtp 的 MTP 对应版本
  • main 分支对比,确认 IS_AGENTIC/DI 路由拆分逻辑及路由就绪校验(wait_or_die)未出现意外回归
  • CI:e2e Test - dsv4-fp4-mi355x-sglang-disagg-agentic-hicache...-hicache-mtp(多节点,mi355x-amds)

Made with Cursor

ichbinblau and others added 7 commits July 23, 2026 06:45
…fill cudagraph for DSv4 agentic disagg

Co-authored-by: Cursor <cursoragent@cursor.com>
…x router policy env var/name

Co-authored-by: Cursor <cursoragent@cursor.com>
…ic-hicache

Resolves review comment: the image bump (v0.5.14 -> v0.5.15.post1),
re-enabled TP8/EP8/DPA hicache arm, and HiCache sizing/NCCL/router
fixes need a changelog entry to trigger the benchmark sweep on merge.

中文:为 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache 添加 perf-changelog.yaml
条目。解决评审意见:镶像升级(v0.5.14 -> v0.5.15.post1)、重新启用的
TP8/EP8/DPA hicache 分支,以及 HiCache 容量/NCCL/路由策略修复,需要一条
changelog 条目才能在合并时触发基准测试扫描。

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Theresa Shan <theresa.shan@amd.com>
Signed-off-by: Theresa Shan <theresa.shan@amd.com>
…entX

DeepSeek-V4's MTP head is a native NextN/EAGLE draft (shipped with the
model), so SGLang defaults the draft path to the main checkpoint and
does not need --speculative-draft-model-path. Adds mtp_flags matching
the upstream SGLang DeepSeek-V4 AMD recipe (EAGLE, topk=1); num-steps
and num-draft-tokens are already derived from DECODE_MTP_SIZE in
server_sglang.sh.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Theresa Shan <theresa.shan@amd.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

…ic-hicache-mtp

Required per AGENTS.md's "Adding a benchmark configuration" rule so the
new EAGLE/MTP arm actually triggers a benchmark sweep on merge.

中文:为 dsv4-fp4-mi355x-sglang-disagg-agentic-hicache-mtp 添加
perf-changelog.yaml 条目。根据 AGENTS.md 的"添加基准测试配置"规则,需要此条目
才能在合并时为新的 EAGLE/MTP 分支触发基准测试扫描。

Co-authored-by: Cursor <cursoragent@cursor.com>
Comment thread configs/amd-master.yaml
Comment thread configs/amd-master.yaml
@ichbinblau ichbinblau added AMD agentx AgentX benchmarks, recipes, and infrastructure labels Jul 23, 2026
@ichbinblau
ichbinblau changed the base branch from amd/agentx-v1.0-th-hicon-dpa to main July 23, 2026 12:02
@ichbinblau
ichbinblau requested a review from a team July 23, 2026 12:02
@ichbinblau ichbinblau changed the title [AMD][AgentX] DSv4 MI355X agentic disagg: add EAGLE/MTP speculative-decoding arm [AMD][AgentX] DSv4 MI355X agentic disagg: add EAGLE/MTP speculative-decoding arm / DSv4 MI355X 智能体分离式推理:新增 EAGLE/MTP 投机解码分支 Jul 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

Signed-off-by: Theresa Shan <theresa.shan@amd.com>
@github-actions

Copy link
Copy Markdown
Contributor

SGLang's PD-disaggregated KV/state transfer requires both ends of the
connection to register the same number of state descriptors. Gating
--speculative-* flags to the decode server only left prefill without the
extra MTP/EAGLE draft state component, causing decode's PD state transfer
to fail with "state component count mismatch (local=3, remote=4)" as soon
as a transfer was attempted on the new disagg-hicache-mtp arm.

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

2 similar comments
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@functionstackx functionstackx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for this PR @ichbinblau ! this is very exciting

1 small nit: per https://github.com/SemiAnalysisAI/InferenceX/blob/main/docs/PR_REVIEW_CHECKLIST.md in order to do fair MTP

SGLANG_SIMULATE_ACC_LEN: <AL from yaml>
SGLANG_SIMULATE_ACC_METHOD: match-expected
 SGLANG_SIMULATE_ACC_TOKEN_MODE: real-draft-token

https://github.com/SemiAnalysisAI/InferenceX/tree/main/golden_al_distribution

functionstackx added a commit that referenced this pull request Jul 25, 2026
…bs (#2330)

The golden AL README only showed vLLM's synthetic_acceptance_length.
Add the SGLang env vars (SGLANG_SIMULATE_ACC_LEN / _METHOD / _TOKEN_MODE)
and TRT-LLM's TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS, including the
off-by-one: the TRT-LLM value counts accepted draft tokens only and
excludes the bonus token, so it must be set to golden AL minus 1
(fractional allowed, e.g. AL 3.5 -> 2.5).

Also teach the codeowner sign-off verify prompt (Check 10) to accept
TRT-LLM configs and to compare the pinned TRT-LLM value plus 1 against
the golden curve. Keep README_zh.md in parity.

Requested in
#2309 (review)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
ichbinblau and others added 3 commits July 26, 2026 14:16
Addresses PR #2309 review feedback (functionstackx):
#2309 (review)

Agentic trace replay doesn't reproduce real token-by-token traffic, so
measured MTP/EAGLE acceptance there isn't representative of production.
Per the AgentX fairness guidelines, wire SGLANG_SIMULATE_ACC_LEN (pinned
to the committed golden AL in golden_al_distribution/dsv4_mtp.yaml,
thinking_on, num_speculative_tokens=3 -> 2.49), SGLANG_SIMULATE_ACC_METHOD
and SGLANG_SIMULATE_ACC_TOKEN_MODE into the decode launch command for
agentic MTP runs, matching the existing GB300 dynamo-sglang DSv4 MTP
recipes. Falls back to real acceptance with a warning if a future draft
length has no committed golden AL wired in yet.

Co-authored-by: Cursor <cursoragent@cursor.com>
Extend the simulated-acceptance lookup to cover DECODE_MTP_SIZE=1 and =2
using the committed thinking_on curve in golden_al_distribution/dsv4_mtp.yaml
(1 -> 1.79, 2 -> 2.27), matching the existing size=3 -> 2.49 entry. Sizes
without a wired golden AL still fall back to real acceptance with a warning.

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions

Copy link
Copy Markdown
Contributor

@ichbinblau

Copy link
Copy Markdown
Collaborator Author

thanks for this PR @ichbinblau ! this is very exciting

1 small nit: per https://github.com/SemiAnalysisAI/InferenceX/blob/main/docs/PR_REVIEW_CHECKLIST.md in order to do fair MTP

SGLANG_SIMULATE_ACC_LEN: <AL from yaml>
SGLANG_SIMULATE_ACC_METHOD: match-expected
 SGLANG_SIMULATE_ACC_TOKEN_MODE: real-draft-token

https://github.com/SemiAnalysisAI/InferenceX/tree/main/golden_al_distribution

Thanks for flagging this @functionstackx — addressed in 5f2bfef and 1f4de33:

@github-actions

Copy link
Copy Markdown
Contributor

3 similar comments
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure AMD full-sweep-enabled

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants